Papers with plug-and-play solution

12 papers
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs (2025.acl-short)

Copied to clipboard

Challenge: Recent advances in Video Large Language Models (Video-LLMs) have achieved exceptional performance on tasks like video question answering and captioning.
Approach: They propose a decoding strategy that leverages sparse top-K attention and dense full attention to accelerate Video-LLMs without loss.
Outcome: The proposed approach achieves a 1.94 walltime speedup in video processing.
Reasoning in Flux: Enhancing Large Language Models Reasoning through Uncertainty-aware Adaptive Guidance (2024.acl-long)

Copied to clipboard

Challenge: Extensive experiments across various reasoning tasks demonstrate that UAG not only enhances the reasoning abilities of LLMs but consistently outperforms several strong baselines with minimal computational overhead.
Approach: They propose an approach to guide LLMs onto an accurate and reliable trajectory by identifying and adjusting uncertainty signals within each step of the reasoning chain.
Outcome: The proposed approach outperforms strong baselines and outperformed strong models with minimal computational overhead.
MPO: Boosting LLM Agents with Meta Plan Optimization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for interactive planning tasks suffer from planning hallucinations and require retraining for each new agent.
Approach: They propose a framework that leverages explicit guidance through meta plans to assist agent planning and enables continuous optimization based on feedback from the agent’s task execution.
Outcome: The proposed framework outperforms existing baselines on two representative tasks and significantly improves task completion efficiency and generalization capabilities.
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on enhancing multi-scale clip representations but lack robust data alignment . inherent data uncertainty renders PRVR vulnerable to distractor videos with spurious similarities .
Approach: proposed framework for partially relevant video retrieval aims to retrieve untrimmed videos partially relevant to a given query.
Outcome: The proposed framework can be seamlessly integrated into existing architectures.
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, but introduce a critical security vulnerability: Knowledge Base Leakage.
Approach: They propose a runtime defense mechanism inspired by stack canaries in software security . canaryRAG embeds carefully designed canary tokens into retrieved chunks and reformulates RAG extraction defense as a dual-path runtime integrity game.
Outcome: The proposed system can detect and prevent RAG Knowledge Base Leakage in real time . it can be integrated into arbitrary RAG pipelines without retraining or structural modifications .
Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiency (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large reasoning models often introduce significant overthinking . this leads to verbose and redundant outputs that hinder efficiency.
Approach: They propose a plug-and-play solution that disables explicit self-reflection . it suppresses tokens such as "Wait" and "Hmm" during inference .
Outcome: The proposed approach reduces chain-of-thought trajectory length by up to 27%–51% in five R1-style model series without compromising model utility.
OjaKV: Context-Aware Online Low-Rank KV Cache Compression (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for inference use static, offline-learned subspaces that perform poorly under distribution shifts.
Approach: They propose a framework that integrates a storage policy with an online subspace adaptation to preserve key-value tokens in full rank as high-fidelity anchors.
Outcome: Experiments show that OjaKV maintains or improves zero-shot accuracy at high compression ratios, achieving the strongest gains on long-context benchmarks requiring complex reasoning.
Refiner: Restructure Retrieved Content Efficiently to Advance Question-Answering Capabilities (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are limited by their parametric knowledge, leading to hallucinations in knowledge-extensive tasks.
Approach: They propose an end-to-end extract-and-restructure paradigm that leverages a single decoder-only LLM to adaptively extract query-relevant contents verbatim along with the necessary context.
Outcome: Experiments show that a trained Refiner outperforms state-of-the-art RAG and compressing approaches in multiple tasks.
Invocation Refiner: A Plug-and-Play Module for Rectifying LLM Tool Invocations (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in Tool-Integrated Reasoning (TIR) however, the practical application is often hindered by frequent errors in tool invocations, such as incorrect tool names, invalid parameters, wrong tool-call order, or malformed invocation formats.
Approach: They propose a specialized post-processing module that performs independent reasoning on the input of a frozen upstream LLM and an advanced RL algorithm to improve the tool-use reliability of base LLMs.
Outcome: The proposed module improves task completion rates and invocation accuracy over the raw outputs of various upstream LLMs on a diverse set of tool-use and reasoning benchmarks.
Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model Merging (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models fail due to performance degradation . existing methods fail for models fine- tuned with low-rank adaptation .
Approach: They propose to constrain the LoRA subspace prior to fine-tuning to ensure that updates relevant to one task do not adversely shift outputs for others.
Outcome: The proposed method can integrate with most existing merging algorithms, reducing unintended interference among tasks.
Structure-aware Fine-tuning for Code Pre-trained Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing CodePTMs are mainly structure-free and structurebased, but how to fine-tune them remains a challenge.
Approach: They propose a plug-and-play fine-tuning method that incorporates structural knowledge into pre-trained code models.
Outcome: The proposed method can benefit CodePTMs more with limited training data.
Accelerated Test-Time Scaling with Model-Free Speculative Sampling (2025.emnlp-main)

Copied to clipboard

Challenge: Language models have demonstrated remarkable capabilities in reasoning tasks through test-time scaling techniques like best-of-N sampling and tree search.
Approach: They propose a model-free speculative decoding approach that exploits redundancy in reasoning trajectories to achieve significant acceleration without compromising accuracy.
Outcome: The proposed approach reduces inference latency by 60-65% while maintaining accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations